Skip to content

pl: mark the unicode translations as verified (t -> T) - #741

Open
michaldziwisz wants to merge 6 commits into
daisy:mainfrom
michaldziwisz:pl-verified-keys
Open

pl: mark the unicode translations as verified (t -> T)#741
michaldziwisz wants to merge 6 commits into
daisy:mainfrom
michaldziwisz:pl-verified-keys

Conversation

@michaldziwisz

Copy link
Copy Markdown
Contributor

Stacked on #738. That PR carries the five substantive fixes; this one is
the mechanical pass. Please merge #738 first, or review this one by looking at
commit 5f1a8dd alone — the other five commits shown here belong to #738 and
will disappear from the diff once it lands.

What this does

Raises 2435 text keys from t/ot/ct to T/OT/CT in
pl/unicode.yaml and pl/unicode-full.yaml. No text is changed.

Why

The Polish rules were written before the lowercase/uppercase convention was in
use, so entries that have been translated all along still carried the
"needs review" key. The audit tool counted every one of them as untranslated:

audit-translations pl, untranslated text:  3162 -> 727

The 3162 made the real gaps invisible. Working out that only ~80 unicode entries
were genuinely untranslated (handled in #738) required writing a separate tool to
compare the text against the English source, because the key case alone says
nothing about whether the work was done.

The remaining 727 live in the rule files. Those need reading rather than a
mechanical pass, so this PR deliberately leaves them alone.

What is not raised

Only entries whose text differs from the English source are touched — that is
the evidence that a translation exists. Entries still holding English text keep
the lowercase key, since raising it would assert a translation that is not there.
Entries whose text legitimately equals English (Roman numerals, ligatures, proper
names, whitespace) were reviewed individually in #738.

Verification

cargo test --test languages Languages::pl    609 passed, 0 failed
diff                                         2436 insertions, 2436 deletions

Zero changes outside key case, checked by normalising the key case in the diff
and confirming no unpaired lines remain:

git diff -U0 | grep -E '^[-+]' | grep -v '^[-+][-+]' \
  | sed -E 's/^.//; s/\b(t|ot|ct|T|OT|CT):/KEY:/' | sort | uniq -c | awk '$1%2!=0'

Without the normalisation 4710 unpaired lines remain, so the check does
discriminate rather than trivially passing.

Kept separate from #738 on purpose: 2435 lines of key case should not bury five
reviewable fixes.

Follow-up to 45aefb2 for Polish. Editors commonly emit U+22A5 where
perpendicularity is meant, so both U+27C2 and U+22A5 now say
"jest prostopadle do". LiteralSpeak keeps reading U+22A5 literally
as "dol".

Previously Polish said "dol" for U+22A5, so the same formula was read
differently in Polish than in English.

Adds up_tack_330 to tests/Languages/pl/alphabets.rs, mirroring the
English test.
Follow-up to 4347d88 for Polish. That commit added the `|` syntax and
updated the Polish tests, but not the Polish rule files, so several
intents lost their function name in speech.

Fixed, with the spoken output before -> after:

  quotient        "podzielone przez z x przecinek, y"
               -> "czesc calkowita z x podzielone przez y"
  remainder       "podzielone przez z x przecinek, y"
               -> "reszta z x podzielone przez y"
  set-difference  "i z wielka a przecinek, wielka b"
               -> "roznica zbiorow z wielka a i wielka b"
  polar-coordinate "przecinek z x przecinek, y"
               -> "wspolrzedna biegunowa z x przecinek, y"

The `|` syntax needs the function-intent rule to call
IntentFunctionUseArityPath / IntentFunctionGlueBefore. Only en and hu
had it, so it is now ported to pl (with "of" -> "z").

Also here:
* SharedRules/geometry.yaml: the `coordinate` rule matched "." instead
  of "not(*[@arg])", so it swallowed coordinate($x,...) intents that
  should fall through to IntentMappings. Its name was also mistranslated
  as "przecinek" (comma) rather than "punkt" (point).
* transpose: fixity order now matches en (postfix first). The
  function form is tested explicitly via intent='transpose:function($x)',
  as en does.
* empty-set: added, it was missing.

NOT adopted: arity templates ("| po | od,do") for sum/product.
intent_function_glue_before in src/infer_intent.rs hardcodes the English
word "of" for the last argument, which yields "suma po i of x" in Polish.
This affects every non-English language; hu avoids it the same way. The
binary separator form works correctly and is what this commit uses.
Follow-up to ec36e05 and 080ca16 for Polish, but it also fixes a
long-standing Polish bug rather than only porting the new code.

navigate.yaml compared the suffix of $NavCommand (always English, e.g.
"ZoomIn") against a substring offset by the length of the SPOKEN $Prefix.
For English those are the same; for any translation they are not. With
"przybliz" (8 chars) vs "Zoom" (4), "ZoomIn" was cut to "" instead of
"In", so ALL 16 branches were dead:

  ZoomIn        -> ""            (want "In")
  ZoomOutAll    -> "ll"          (want "OutAll")
  MoveNext      -> "t"           (want "Next")
  DescribeNext  -> "ibeNext"     (want "Next")

Users heard "przejdz; do mianownika" with no direction, never
"przejdz w prawo". Using $CommandOffset, as ec36e05 introduced, fixes
all of them.

Two more things here:

* Polish needs two verbs where English reuses "zoom": "przybliz na
  zewnatrz" (zoom in outwards) is self-contradictory, so ZoomOut* now
  says "oddal". The direction word for plain In/Out is dropped, as the
  verb already carries it: "przybliz" / "oddal", and
  "przybliz maksymalnie" / "oddal maksymalnie".
* Ports into-or-out-of-prefix-or-silent-without-parts (080ca16) and the
  SpeakIntentName fallbacks in into-or-out-of-default, keeping our own
  Polish preposition logic (including "ze stopnia" euphony).

Adds zoom_speech_pl and move_char_speech_pl. Both fail on the old
formula, showing the missing direction word, so they do discriminate.
Follow-up to fbc49bb (daisy#679) for Polish. The HasVisibleColumnLine /
HasVisibleRowLine rules were missing from pl/SharedRules/default.yaml,
so visible lines in a matrix were silent for Polish users, e.g. an
augmented matrix was read exactly like a plain one.

  before: "2 na 3 macierz rozszerzona; wiersz 1; 3, 1, 4; ..."
  after:  "2 na 3 macierz rozszerzona; wiersz 1; 3, 1, separator, 4; ..."

Two existing tests (augmented_matrix_2x3, augmented_matrix_3x4_end_matrix)
were pinning the pre-daisy#679 output; their English counterparts already
expect "separator", so they are updated rather than worked around.

Adds dashed_augmented_matrix_separator and matrix_row_separator,
ported from tests/Languages/en/mtable.rs. Removing either rule fails
all four tests, so they discriminate.
The audit tool reported 81 unicode entries whose text equals the English
source. Reviewing them one by one, only two were actually untranslated:

  U+2127  "mhos"  -> "mho"    (unit, as nb and sv have it)
  U+2644  "Saturn"           (Polish spelling is the same; key raised)

The other 75 are correct as-is and only needed the verified key:

  * 30 Roman numerals (U+2160..U+217F) spelled out letter by letter
  * 29 space, PUA and zero-width entries with no speech at all
  *  6 typographic ligatures (ff, fl, ffi, ffl, ft, st)
  * 10 proper names and symbols (spesmilos, paragraphos, hypodiastole,
       digamma, differential d, imaginary j, oV, pH)

Four more in unicode.yaml (digit separator, space, U+2062, U+2063)
likewise carry no translatable speech.

Raising the key on an entry whose text legitimately matches English is
what the convention is for; it is not the same as marking English text
as verified. Only the wording of "mhos" changed - the diff is otherwise
key case only, checked line by line.

Audit's "untranslated" for unicode files: 81 -> 0.
Mechanical change: 2435 text keys raised from t/ot/ct to T/OT/CT in the
two unicode files. No text is touched.

The Polish rules were written before the lowercase/uppercase convention
was in use, so entries that have been translated all along still carried
the "needs review" key. The audit tool therefore reported them as
untranslated, which made the real gaps impossible to see:

  audit-translations pl, untranslated text:  3162 -> 727

The remaining 727 are in the rule files, where the entries need reading
rather than a mechanical pass, so they are deliberately left alone.

Every raised entry was checked to differ from the English source, i.e.
it really is translated. Entries whose text legitimately equals English
(Roman numerals, ligatures, proper names, whitespace) were handled
separately in the previous PR; entries still holding English text are
NOT raised, since that would assert a translation that does not exist.

Verification:

  cargo test --test languages Languages::pl   609 passed, 0 failed
  diff: 2436 insertions, 2436 deletions, zero changes outside key case

The last point is checked by normalising the key case in the diff and
confirming no unpaired lines remain (without normalising, 4710 remain,
so the check does discriminate).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Triage

Development

Successfully merging this pull request may close these issues.

1 participant